Europarl corpus について

Words near each other

・ Europaeus
・ EUROPAfest
・ Europafilm
・ Europagymnasium Auhof
・ Europahalle
・ Europahaus
・ Europalestine
・ Europalia
・ Europan
・ Europanto
・ Europapress Holding
・ Europarc
・ EUROPARC Federation
・ EuroPark
・ Europark Idroscalo Milano
・ Europarl corpus
・ EuroparlTV
・ Europasaurus
・ Europass
・ Europass (disambiguation)
・ Europass centro studi europeo
・ Europatat
・ Europaturm
・ Europaviertel
・ Europaviertel (Wiesbaden)
・ Europay International
・ Europcar
・ Europe
・ Europe '51
・ Europe '72

Dictionary Lists

mini英和辞書

翻訳と辞書　辞書検索 [ 開発暫定版 ]

スポンサードリンク

Europarl corpus ：ウィキペディア英語版

Europarl corpus
The Europarl Corpus is a corpus (set of documents) that consists of the proceedings of the European Parliament from 1996 to the present.
In its first release in 2001, it covered eleven official languages of the European Union (Danish, Dutch, English, Finnish, French, German, Greek, Italian, Portuguese, Spanish, and Swedish).〔 With the political expansion of the EU the official languages of the ten new member states have been added to the corpus data.〔 The latest release (2012)〔 comprised up to 50 million words per language with the newly added languages being slightly underrepresented as data for them is only available from 2007 onwards.〔
The data that makes up the corpus was extracted from the website of the European Parliament and then prepared for linguistic research.〔 After sentence splitting and tokenization the sentences were aligned across languages with the help of an algorithm developed by Gale & Church (1993).〔
The corpus has been compiled and expanded by a group of researchers led by Philipp Koehn at Edinburgh University. Initially it was designed for research purposes in statistical machine translation (SMT). However, since its first release it has been used for multiple other research purposes, including for example word sense disambiguation.
== Europarl Corpus and Statistical Machine Translation ==
In his paper “Europarl: A Parallel Corpus for Statistical Machine Translation” (2005) Koehn sums up in how far the Europarl corpus is useful for research in SMT. He uses the corpus to develop SMT systems translating each language into each of the other ten languages of the corpus making it 110 systems. This enables Koehn to establish SMT systems for uncommon language pairs that have not been considered by SMT developers beforehand, such as Finnish-Italian for example.

抄文引用元・出典: フリー百科事典『ウィキペディア（Wikipedia）』
■ウィキペディアで「Europarl corpus」の詳細全文を読む

スポンサードリンク

翻訳と辞書 : 翻訳のためのインターネットリソース